Papers with LLaMA 3.1
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)
Copied to clipboard
| Challenge: | Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology. |
| Approach: | They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology. |
| Outcome: | The results challenge the validity of current benchmark-based claims about social reasoning in large language models. |
Rhetorical Device-Aware Sarcasm Detection with Counterfactual Data Augmentation (2025.findings-acl)
Copied to clipboard
| Challenge: | Sarcasm is a complex form of sentiment expression widely used in human daily life. |
| Approach: | They propose a device-aware sarcasm dataset with counterfactually augmented data to capture its complexity. |
| Outcome: | The proposed dataset shows that it is more balanced than zero-shot models. |